Genome analysis with inter-nucleotide distances

نویسندگان

  • Vera Afreixo
  • Carlos A. C. Bastos
  • Armando J. Pinho
  • Sara P. Garcia
  • Paulo Jorge S. G. Ferreira
چکیده

MOTIVATION DNA sequences can be represented by sequences of four symbols, but it is often useful to convert the symbols into real or complex numbers for further analysis. Several mapping schemes have been used in the past, but they seem unrelated to any intrinsic characteristic of DNA. The objective of this work was to find a mapping scheme directly related to DNA characteristics and that would be useful in discriminating between different species. Mathematical models to explore DNA correlation structures may contribute to a better knowledge of the DNA and to find a concise DNA description. RESULTS We developed a methodology to process DNA sequences based on inter-nucleotide distances. Our main contribution is a method to obtain genomic signatures for complete genomes, based on the inter-nucleotide distances, that are able to discriminate between different species. Using these signatures and hierarchical clustering, it is possible to build phylogenetic trees. Phylogenetic trees lead to genome differentiation and allow the inference of phylogenetic relations. The phylogenetic trees generated in this work display related species close to each other, suggesting that the inter-nucleotide distances are able to capture essential information about the genomes. To create the genomic signature, we construct a vector which describes the inter-nucleotide distance distribution of a complete genome and compare it with the reference distance distribution, which is the distribution of a sequence where the nucleotides are placed randomly and independently. It is the residual or relative error between the data and the reference distribution that is used to compare the DNA sequences of different organisms.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Sequencing and Molecular Analysis of ATP 6 and ATP 8 of Mitochondrial Genome in Khorasanian Native Chickens

In order to perform breeding programs and improve production of native chickens, preserving genetic diversity in different areas of Iran is important due to the reduced available population. Genome sequencing is considered the most functional approach to determine the phylogeny relation between close populations. The aim of the present study was the evaluation of the phylogeny and genetic nucle...

متن کامل

Inter-dinucleotide distances in the human genome: an analysis of the whole-genome and protein-coding distributions

We study the inter-dinucleotide distance distributions in the human genome, both in the whole-genome and protein-coding regions. The inter-dinucleotide distance is defined as the distance to the next occurrence of the same dinucleotide. We consider the 16 sequences of inter-dinucleotide distances and two reading frames. Our results show a period-3 oscillation in the protein-coding inter-dinucle...

متن کامل

DNA Polymorphisms at Candidate Gene Loci and Their Relation with Milk Production Traits in Murrah Buffalo (Bubalus bubalis)

DNA polymorphism within diacylglycerol transferase 2 (DGAT2) / monoacyl glycerol transferases 2 (MOGAT2), leptin and butyrophilin genes were analysed using PCR-SSCP in Murrah buffalo. The single strand conformation polymorphism (SSCP) analysis of amplified gene fragment in exon 5 of MOGAT2, exon 3 of leptin and intron 1 of butyrophilin gene revealed different patterns. A, B and C showed the fol...

متن کامل

Comparison of Phylogenetic and Evolutionary of Nucleotide Squences of HVR1 region of Mitochondria genom in Goats and Other Livestock Species

     Maintaining genomic diversity in goat populations in different parts of Iran is essential for breeding programs, increasing production, survival, resistance to diseases, and various environmental changing conditions. The aim of the present study was to determine the sequence of HVR1 from the mitochondrial genome of Iranian native goats including Sistani, Pakistani, Black and Lorry ecotypes...

متن کامل

BAC-End Microsatellites from Intra and Inter-Genic Regions of the Common Bean Genome and Their Correlation with Cytogenetic Features

Highly polymorphic markers such as simple sequence repeats (SSRs) or microsatellites are very useful for genetic mapping. In this study novel SSRs were identified in BAC-end sequences (BES) from non-contigged, non-overlapping bacterial artificial clones (BACs) in common bean (Phaseolus vulgaris L.). These so called "singleton" BACs were from the G19833 Andean gene pool physical map and the new ...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره 25  شماره 

صفحات  -

تاریخ انتشار 2009